Training full-duplex LLMs for spoken dialogue
How are full-duplex (simultaneous bidirectional speech) LLMs trained, and which architectures, data, and training methods achieve low-latency, interruptible spoken dialogue?
Full-duplex spoken dialogue models — systems that listen and speak at the same time — became a distinct research field between 2023 and 2026, moving from the first speech-in/speech-out language models to a crowded design space of dual-stream transformers, frozen-LLM adapters, and single-stream 'native duplex' recipes. The evidence shows training data that contains genuine overlapping speech is the binding constraint, not architecture; latency and interruption handling are now measurable but benchmarks disagree about what to measure; and post-training alignment (RL/DPO on turn-taking behavior) is the fastest-moving lever. Confidence is moderate: the field is young, most frontier systems are preprints, and no standardized evaluation exists yet.
Updated 7 Aug 2026111 sources2023–2026Deep23 min read
full-duplex speech · spoken dialogue systems · speech language models · turn-taking · interruption handling · audio codecs